挑最差的 3 號出來研究
3. Bypasses stalled add 108 107 92 109
case 3: independent branch, taken
addi x10, x0, 80
addi x11, x0, 2
div x1, x10, x11 # x1 = 40
add x2, x1, x1 # waits for DIV; x2 = 80
beq x0, x0, target # ready without DIV
addi x2, x0, 99 # skipped architecturally
target: mul x3, x10, x11 # x3 = 160
分析的部分趕工中,晚點補上!
另外加了一個 sequence 讓有 rename recovery 的 out of order cpu 平反一下
addi x10, x0, 120
addi x11, x0, 3
div x1, x10, x11 # x1 = 40
div x2, x1, x11 # Waits for first DIV; x2 = 13
add x3, x2, x2 # Waits for second DIV; x3 = 26
beq x1, x0, done # Needs only first DIV; NOT taken
# Independent of both divides: a chain of useful younger work
addi x4, x0, 7
addi x5, x4, 1
addi x6, x5, 1
addi x7, x6, 1
addi x8, x7, 1
addi x9, x8, 1
addi x10, x9, 1
addi x11, x10, 1
addi x12, x11, 1
addi x13, x12, 1
addi x14, x13, 1
done:
add x20, x14, x3 # Joins both dependency chains; x20 = 43